Ashley writes about Needle 2, a 14MB function-calling LLM from Cactus Compute that converts plain English prompts into local actions on a Raspberry Pi 5 using CPU alone. Rather than acting as a general chatbot, the model is purpose-built to select from declared Python functions and fill in their arguments, running entirely offline after a one-time download. In benchmark runs, inference latency ranges from 76 to 149 milliseconds, and the model correctly refuses questions outside its declared tool set.
- Native session is ~28MB; the full Python process peaks at 43–46.4MB
- Weights and code released under Apache 2.0 on Hugging Face and GitHub
- Can be fine-tuned locally on a laptop for a specific set of tools
- Eben Upton's endorsement: "Needle 2 is rather excellent"